Papers with Multimodal foundation models

4 papers
Fact-Aware Multimodal Retrieval Augmentation for Accurate Medical Radiology Report Generation (2025.naacl-long)

Copied to clipboard

Challenge: Existing multimodal foundation models suffer from serious factual inaccuracy in radiology report generation.
Approach: They propose a fact-aware multimodal retrieval-augmented pipeline for generating accurate radiology reports using RadGraph.
Outcome: The proposed multimodal retrieval-augmented pipeline outperforms state-of-the-art retrievers on language generation and radiology-specific metrics.
How do Multimodal Foundation Models Encode Text and Speech? An Analysis of Cross-Lingual and Cross-Modal Representations (2025.naacl-short)

Copied to clipboard

Challenge: Recent advances in foundation models have sparked growing interest in expanding their text processing capabilities to speech.
Approach: They analyze the model activations from semantically equivalent sentences across languages in the text and speech modalities and examine how text and spoken are represented in recent multimodal foundation models.
Outcome: The proposed models exhibit cross-lingual differences, but are not explicitly trained for modality-agnostic representations.
Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding (2025.findings-naacl)

Copied to clipboard

Challenge: Multimodal foundation models have demonstrated significant success in tasks such as visual captioning, question answering, and image-text retrieval.
Approach: They propose a specialized cognitive module, temporal working memory, which selectively retains task-relevant information across temporal dimensions.
Outcome: The module retains task-relevant information across temporal dimensions, ensuring that critical details are preserved throughout the processing of video and audio content.
SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in multimodal large language models have increased their vulnerability to adversarial manipulation.
Approach: They propose to target audio-only adversarial attacks on multimodal audio–video–language models . they show that attacks can be successful at low perceptual distortions .
Outcome: The proposed models achieve up to 96% success rate under realistic conditions . the proposed models are more robust to noise than to noise and distortion than to speech recognition systems .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations